ESC

Type to search articles...

No articles found.

↑ ↓ Navigate ↵ Open
Esc Close
Blog Tags GitHub
All tags

LLM Serving

1 post tagged with "LLM Serving"

July 23, 2026

Robust KV Cache Management for LLM Serving under Output Token Length Uncertainty

提出了一种基于Wasserstein DRO的鲁棒KV cache管理框架,联合优化GPU并行配置、KV cache预留、请求路由和前缀缓存,在输出token长度不确定性下实现自适应内存分配和尾延迟控制。核心贡献包括临界分位数结构理论证明、BCD-DRO分解算法和滚动时域自适应策略。

LLM Serving KV Cache Distributionally Robust Optimization Wasserstein DRO GPU Resource Management Queueing Theory LLM Inference

Navigation

  • Work

Resources

  • Lexington Themes.

Socials

  • @Mike_Andreuzza
© 2025 MicroStudio. All rights reserved.

MicroStudio is not affiliated with Stripe, Breeew, Astro, or Tailwind Labs, nor is it endorsed or sponsored by them.